Papers with Social Media

6 papers
SEDTWik: Segmentation-based Event Detection from Tweets Using Wikipedia (N19-3)

Copied to clipboard

Challenge: Recent work on event detection from tweets has focused on localized events or breaking news only.
Approach: They propose to split tweets into segments, extract bursty segments, cluster them, summarize them.
Outcome: The proposed system can detect newsworthy events occurring at different locations of the world from a wide range of categories.
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (2026.findings-eacl)

Copied to clipboard

Challenge: Social media platforms such as X (formerly Twitter), Facebook, and Reddit generate user-generated content.
Approach: They propose a framework to assess privacy risks in social media by evaluating vulnerabilities across six dimensions: data collection, preprocessing, visibility, fairness, computational risk, and regulatory compliance.
Outcome: The proposed framework assesses privacy risks across six dimensions . it achieves F1-scores of 0.58–0.84, but incurs 1% - 23% drop under fine-tuning .
Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection (2021.acl-short)

Copied to clipboard

Challenge: a lack of labeled, non-English resources for hate speech detection limits research on hate speech . a recent study shows that zero-shot, cross-lingual learning models cannot be used as they are . lack of consistency limits research, and lack of models for non-english languages limits learning .
Approach: They propose a zero-shot, cross-lingual transfer learning framework for hate speech detection . they use benchmark data sets in English, Italian, and Spanish to detect hate speech .
Outcome: The proposed framework can't be used as it is, but needs to be carefully designed, the authors say . they find that non-hateful, language-specific taboo interjections are misinterpreted as signals of hate speech .
MediaHG: Rethinking Eye-catchy Features in Social Media Headline Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Creating a good headline on social media platforms requires a disentanglement-based model to balance the content and contextual features.
Approach: They propose a disentanglement-based headline generation model which can balance the content and contextual features by incorporating contrastive learning and auxiliary multi-tasking to choose the best domain-suitable headline.
Outcome: The proposed model can balance content and contextual features, while allowing bloggers to obtain more site traffic and profits while readers can have easier access to topics of interest.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis.
Approach: They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them.
Outcome: The proposed package is the first to be released in Bulgarian and is available for free.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations